This archive contains the data relevant to the sequencing analysis, either with similarity clustering or without. The columns inside the sequencing_numbers tables describe the quality threshold a sequence had to reach to not be discarded (either 10, 20 or 30), followed by the pandaSEQ quality the assembled sequence had to reach (either 0.3, 0.6 or 0.9) and the percentage of reads from the raw sequencing results used. The other rows describe, for a given file and combination of minimal PANDAseq and mean quality score, the lowest percentage of the raw reads required for successful decoding and the highest evaluated percentage of raw reads, for which the decoding failed.

It further contains the results for the rate analysis, including final parameters, the results of the error-correction performance evaluations, the encoded data for each code, split into 10mers for the mCGR evaluation, a table containing the parameters used for the cost analysis for each code and a config file for the error simulator MESA (https://mesa.mosla.de/) that contains the used parameter combination for simulating the DNA storage channel under realistic conditions.
